
Allen Institute for AI released Olmo-core 3, an open-source framework designed to train mixture-of-experts models up to a trillion parameters. It abandons fully sharded data parallelism for a distributed data parallelism setup that keeps experts resident on the GPUs, routing data directly to avoid repeated weight gathering. In benchmarks on NVIDIA B300 hardware, a 47-billion-parameter MoE hit 52,000 tokens per second per GPU, a 2.7x improvement over their older stack.
It is worth a look if you are building massive sparse models and need an alternative to Megatron-Core, especially with the added MXFP8 support yielding a 21 percent throughput bump over BF16. However, the complexity here is substantial, meaning it is entirely irrelevant if you are sticking to dense architectures or smaller academic setups.
Leave a comment